NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

The Quantum Tortoise and the Classical Hare: When Will Quantum Computers Outpace Classical Ones and When Will They Be Left Behind?

https://doi.org/10.1109/JPROC.2025.3574102

CHOI, SUKWOONG; MOSES, WILLIAM S; THOMPSON, NEIL (June 2025, Proceedings of the IEEE)

Free, publicly-accessible full text available June 19, 2026
The MLIR Transform Dialect: Your Compiler Is More Powerful Than You Think

https://doi.org/10.1145/3696443.3708922

Lücke, Martin Paul; Zinenko, Oleksandr; Moses, William S; Steuwer, Michel; Cohen, Albert (March 2025, ACM)

Free, publicly-accessible full text available March 1, 2026
Distributed MIS in O(log log n) Awake Complexity

https://doi.org/10.1145/3583668.3594574

Dufoulon, Fabien; Moses, William K.; Pandurangan, Gopal (June 2023, PODC '23: Proceedings of the 2023 ACM Symposium on Principles of Distributed Computing)

Full Text Available
Transparent Checkpointing for Automatic Differentiation of Program Loops Through Expression Transformations

https://doi.org/10.1007/978-3-031-36024-4_37

Schanen, Michel; Narayanan, Sri Hari; Williamson, Sarah; Churavy, Valentin; Moses, William S; Paehler, Ludger. (June 2023, Computational Science – ICCS 2023)

Automatic differentiation (AutoDiff) in machine learning is largely restricted to expressions used for neural networks (NN), with the depth rarely exceeding a few tens of layers. Compared to NN, numerical simulations typically involve iterative algorithms like time steppers that lead to millions of iterations. Even for modest-sized models, this may yield infeasible memory requirements when applying the adjoint method, also called backpropagation, to time-dependent problems. In this situation, checkpointing algorithms provide a trade-off between recomputation and storage. This paper presents the package Checkpointing.jl that leverages expression transformations in the programming language Julia and the package ChainRules.jl to automatically and transparently transform loop iterations into differentiated loops. The user may choose between various checkpointing algorithm schemes and storage devices. We describe the unique design of Checkpointing.jl and demonstrate its features on an automatically differentiated MPI implementation of Burgers’ equation on the Polaris cluster at the Argonne Leadership Computing Facility.
more » « less
Full Text Available
High-Performance GPU-to-CPU Transpilation and Optimization via High-Level Parallel Constructs

https://doi.org/10.1145/3572848.3577475

Moses, William S; Ivanov, Ivan R; Domke, Jens; Endo, Toshio; Doerfert, Johannes; Zinenko, Oleksandr (February 2023, ACM)
NA (Ed.)
While parallelism remains the main source of performance,architectural implementations and programming modelschange with each new hardware generation, often leadingto costly application re-engineering. Most tools for perfor-mance portability require manual and costly application port-ing to yet another programming model.We propose an alternative approach that automaticallytranslates programs written in one programming model(CUDA), into another (CPU threads) based on Polygeist/MLIR.Our approach includes a representation of parallel constructsthat allows conventional compiler transformations to ap-ply transparently and without modification a nd enablesparallelism-specific optimizations. We evaluate our frame-work by transpiling and optimizing the CUDA Rodinia bench-mark suite for a multi-core CPU and achieve a 58% geomeanspeedup over handwritten OpenMP code. Further, we showhow CUDA kernels from PyTorch can efficiently run andscale on the CPU-only Supercomputer Fugaku without userintervention. Our PyTorch compatibility layer making use oftranspiled CUDA PyTorch kernels outperforms the PyTorchCPU native backend by 2.7×.
more » « less
Full Text Available
Brief Announcement: Distributed MST Computation in the Sleeping Model: Awake-Optimal Algorithms and Lower Bounds

https://doi.org/10.1145/3519270.3538459

Augustine, John; Moses, William K.; Pandurangan, Gopal (July 2022, PODC'22: Proceedings of the 2022 ACM Symposium on Principles of Distributed Computing)

Full Text Available
Scalable Automatic Differentiation of Multiple Parallel Paradigms through Compiler Augmentation

https://doi.org/10.1109/SC41404.2022.00065

Moses, William S.; Narayanan, Sri Hari; Paehler, Ludger; Churavy, Valentin; Schanen, Michel; Hückelheim, Jan; Doerfert, Johannes; Hovland, Paul (November 2022, IEEE)

Full Text Available
Singularly Near Optimal Leader Election in Asynchronous Networks

https://doi.org/10.4230/LIPIcs.DISC.2021.27

Kutten, Shay; Moses, William K.; Pandurangan, Gopal; Peleg, David (January 2021, Proceedings of the International Symposium on Distributed Computing (DISC) 2021)
Gilbert, Seth (Ed.)
This paper concerns designing distributed algorithms that are singularly optimal, i.e., algorithms that are simultaneously time and message optimal, for the fundamental leader election problem in asynchronous networks. Kutten et al. (JACM 2015) presented a singularly near optimal randomized leader election algorithm for general synchronous networks that ran in O(D) time and used O(m log n) messages (where D, m, and n are the network’s diameter, number of edges and number of nodes, respectively) with high probability. Both bounds are near optimal (up to a logarithmic factor), since Ω(D) and Ω(m) are the respective lower bounds for time and messages for leader election even for synchronous networks and even for (Monte-Carlo) randomized algorithms. On the other hand, for general asynchronous networks, leader election algorithms are only known that are either time or message optimal, but not both. Kutten et al. (DISC 2020) presented a randomized asynchronous leader election algorithm that is singularly near optimal for complete networks, but left open the problem for general networks. This paper shows that singularly near optimal (up to polylogarithmic factors) bounds can be achieved for general asynchronous networks. We present a randomized singularly near optimal leader election algorithm that runs in O(D + log² n) time and O(m log² n) messages with high probability. Our result is the first known distributed leader election algorithm for asynchronous networks that is near optimal with respect to both time and message complexity and improves over a long line of results including the classical results of Gallager et al. (ACM TOPLAS, 1983), Peleg (JPDC, 1989), and Awerbuch (STOC, 89).
more » « less
Full Text Available
Instead of Rewriting Foreign Code for Machine Learning, Automatically Synthesize Fast Gradients

Moses, William; Churavy, Valentin (January 2020, Advances in neural information processing systems)

Applying differentiable programming techniques and machine learning algorithms to foreign programs requires developers to either rewrite their code in a machine learning framework, or otherwise provide derivatives of the foreign code. This paper presents Enzyme, a high-performance automatic differentiation (AD) compiler plugin for the LLVM compiler framework capable of synthesizing gradients of statically analyzable programs expressed in the LLVM intermediate representation (IR). Enzyme synthesizes gradients for programs written in any language whose compiler targets LLVM IR including C, C++, Fortran, Julia, Rust, Swift, MLIR, etc., thereby providing native AD capabilities in these languages. Unlike traditional source-to-source and operator-overloading tools, Enzyme performs AD on optimized IR. On a machine-learning focused benchmark suite including Microsoft's ADBench, AD on optimized IR achieves a geometric mean speedup of 4.2 times over AD on IR before optimization allowing Enzyme to achieve state-of-the-art performance. Packaging Enzyme for PyTorch and TensorFlow provides convenient access to gradients of foreign code with state-of-the-art performance, enabling foreign code to be directly incorporated into existing machine learning workflows.
more » « less
Full Text Available
Reverse-mode automatic differentiation and optimization of GPU kernels via enzyme

https://doi.org/10.1145/3458817.3476165

Moses, William S.; Churavy, Valentin; Paehler, Ludger; Hückelheim, Jan; Narayanan, Sri Hari; Schanen, Michel; Doerfert, Johannes (November 2021, SC '21: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis)

Computing derivatives is key to many algorithms in scientific computing and machine learning such as optimization, uncertainty quantification, and stability analysis. Enzyme is a LLVM compiler plugin that performs reverse-mode automatic differentiation (AD) and thus generates high performance gradients of programs in languages including C/C++, Fortran, Julia, and Rust. Prior to this work, Enzyme and other AD tools were not capable of generating gradients of GPU kernels. Our paper presents a combination of novel techniques that make Enzyme the first fully automatic reversemode AD tool to generate gradients of GPU kernels. Since unlike other tools Enzyme performs automatic differentiation within a general-purpose compiler, we are able to introduce several novel GPU and AD-specific optimizations. To show the generality and efficiency of our approach, we compute gradients of five GPU-based HPC applications, executed on NVIDIA and AMD GPUs. All benchmarks run within an order of magnitude of the original program's execution time. Without GPU and AD-specific optimizations, gradients of GPU kernels either fail to run from a lack of resources or have infeasible overhead. Finally, we demonstrate that increasing the problem size by either increasing the number of threads or increasing the work per thread, does not substantially impact the overhead from differentiation.
more » « less
Full Text Available

Search for: All records